Papers with syntax-based word embeddings

    1 papers
    Building a Web-Scale Dependency-Parsed Corpus from CommonCrawl (L18-1)

    Copied to clipboard

    Challenge: DepCC is the largest-to-date linguistically analyzed corpus in English . large corpora are essential for the modern data-driven approaches to natural language processing .
    Approach: They present a large-to-date linguistically analyzed corpus in English with 365 million documents . they build an index of all sentences and their linguistic meta-data enabling quick search across the corpus .
    Outcome: The proposed model outperforms state-of-the-art models on smaller corpora on the SimVerb3500 dataset.

    What is GenGO?

    GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

    Information

    About
    Limitations